Conversation
… selector
The xPyD harness already models pools that span multiple nodes (xP/yD count
NODES, roles are MASTER/CHILD via --data-parallel-start-rank + --headless), but
the wideEP path hardcoded TP=1: dp_size was xP*GPUS_PER_NODE and the connector
emitted a literal `-tp 1`. That makes any model whose REPLICATED (non-expert)
weights exceed one GPU inexpressible - e.g. Kimi-K3 on MI300X, where TP1/DP16
needs 190.7 GiB/GPU (> 192 GB HBM) and TP2 is required to fit.
- vllm_disagg.sh: topology math is TP-aware. dp_per_node = GPUS_PER_NODE/TP_SIZE
feeds dp_size, dp_size_local and both start-rank calculations. TP_SIZE defaults
to 1 and is validated to divide GPUS_PER_NODE.
- moriio.sh: emit -tp ${TP_SIZE}; clamp --api-server-count to dp_size (it must be
<= data-parallel-size or the frontend DP balancer routes to ranks that do not
exist - only reachable once TP>1 shrinks dp_size below GPUS_PER_NODE).
- moriio.sh: advertise the peer pool's node IPs as moriio_pod_hosts when a pool
spans >1 node. Without it the connector falls back to the peer MASTER only and
KV writes aimed at ranks on a peer CHILD node silently miss, so those ranks
decode with no context. Emitted only when xP>1 || yD>1.
- run_xPyD_models.slurm: forward TP_SIZE/GPUS_PER_NODE, wire BENCHMARK_SCRIPT=niah
to the already-present benchmark_niah.sh, forward NIAH_* knobs.
Default-neutral: at xP=1 yD=1 TP_SIZE=1 - the topology every registered entry
uses - DRY_RUN output is byte-identical to develop across all four node roles.
At xP=2 yD=2 TP_SIZE=2 it emits tp=2, dp_size=8, dp_local=4, child start-rank 4
(= EP16 per pool over 16 GPUs).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…Kimi-K3 recipe
Two parts.
1) moriio.sh wideEP dropped ${model_args[@]} on the floor. models.yaml documents
per-model dp: tuning as supported ("both connectors now append it") and rixl's
deepep path does append it, but the moriio wideEP branch parsed the flags,
logged them, and then emitted an argv without them. Verified with a sentinel
flag: present in the log line, absent from the command. Now appended, matching
rixl. Every existing wideEP model has an empty dp: block, so emitted argv is
unchanged for all of them.
2) Kimi-K3 recipe, entirely as data: env: block for the gfx942 knobs
(VLLM_ROCM_USE_AITER_MLA=0 - the AITER MLA kernel is gfx950-only -
AITER_SITUV2_A8W4=1 for the packed-int4 SiTUv2 MoE path, KDA conv state layout,
40e9 KV cache for >600K contexts, per-role cudagraph + MoRI backends), and a
dp: block for the serve flags the connector does not emit (reasoning parser,
1M max-model-len, batched-tokens pinned at 2048, MoE quantization-config).
Registered in MORI_EP_VALID_MODELS and WIDE_EP_ONLY_MODELS.
The JSON quantization-config carries its own shell quotes: the launcher
word-splits dp: via eval, which would otherwise strip the JSON's double quotes
and hand vLLM invalid JSON. Verified the emitted value parses as JSON.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docker/pyt_vllm_kimi_k3_mi300x.ubuntu.amd.Dockerfile
The shared vllm_disagg_inference stack with the pins K3 needs on gfx942
(MoRI 1.2.2, AITER 0.1.19 + flydsl 0.2.4, the K3+MoRIIO vLLM commit) plus the
K3-aware AITER graft from the public vendor image - without that graft the K3
MoE profiling shape finds no tuned FlyDSL config, falls back to a heuristic
kernel and aborts LLVM inside determine_available_memory.
Differences from the recipe this is ported from:
- every source pinned to an immutable SHA, not a fork branch name (branches on
personal forks can be force-pushed; MAD needs the image rebuildable later)
- GH_TOKEN build-arg dropped: all three repos are public, and a token passed
this way is recorded in image metadata
- the flydsl 0.2.4 re-pin happens at build time, so the launcher no longer runs
pip install inside every container at serve time
- WITH_NIXL defaults to 0 (K3 uses the moriio connector only)
- no runtime patchers: all connector/KDA fixes are committed in the pinned vLLM
scripts/vllm_multinode/
Colocated (single-instance) multi-node serving - the counterpart to
vllm_dissag, for models that do not fit one node but want lowest single-request
latency rather than disagg throughput. TP within a node, PP across nodes.
It owns only node discovery, the container launch and the head/worker split;
models.yaml, socket_barrier.py, benchmark_xPyD.sh, benchmark_niah.sh and
parse_to_csv.py are reused from vllm_dissag (scripts/ is mounted whole), so the
model recipe and the CSV pipeline have one home.
The three K3 colocated variants collapse to this one script: dry-run confirms
pp2xtp8 / allgather / moriep argv differ only by --enable-expert-parallel and
--all2all-backend, so they become three models.json entries rather than three
near-identical run.sh copies.
Not built here - this environment has no docker and one 8-GPU node; the harness
is verified by DRY_RUN argv inspection only.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
models.json: four entries, all skip_gpu_arch gfx950 (the existing four Kimi-K3 entries are gfx950/TP8 single-node and all carry skip_gpu_arch gfx942, so K3 is currently skipped entirely on MI300X - this is that gap): pyt_vllm_disagg_mori_kimi-k3 -N 4 xP2 yD2 TP2 -> EP16/pool pyt_vllm_kimi-k3_mi300x_pp2xtp8 -N 2 TP8 x PP2, no EP pyt_vllm_kimi-k3_mi300x_wideep_allgather -N 2 + EP, allgather_reducescatter pyt_vllm_kimi-k3_mi300x_wideep_moriep -N 2 + EP, mori_low_latency The three colocated entries differ only in ENABLE_EP / ALL2ALL_BACKEND / AITER_SITUV2_A8W4, which is why they share one launcher. Result reporting: BENCHMARK_SCRIPT=niah previously produced logs and nothing else, so a model declaring multiple_results would have reported no metric. - parse_to_csv.py gains --niah: parses benchmark_niah.py output and writes the same 29-column madengine perf.csv as the throughput path (verified identical column sets). One row per context size, performance = needles found /10. A size that errored is written as a FAILURE row with performance 0 rather than dropped, so a pass->crash regression is visible instead of silent. - benchmark_niah.sh calls it. - Both slurm launchers copy /run_logs/$SLURM_JOB_ID/perf.csv to ./perf_$MODEL_NAME.csv on completion: madengine resolves multiple_results against the job directory it launched from, not the container log mount. No-op for entries that do not declare multiple_results. Verified: madengine discover goes 133 -> 137 with exactly these four added and none removed; every referenced dockerfile and script exists; -N matches each topology; the int4 quantization-config survives shell word-splitting as valid JSON; DeepSeek moriio and rixl/deepep DRY_RUN output is byte-identical to develop. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ports the writeups from PR ROCm#193 next to the existing gfx950 K3 doc, keyed to the madengine tags rather than to shell scripts: topology rationale (why MI300X needs multi-node and why the disagg pool needs TP2), the gfx942 knobs and where each one lives now, the NIAH results, and the three root-cause fixes. Corrections carried in: - the NIAH ceiling is 900K throughout. The upstream READMEs still said 300K in seven places after two later commits raised it to 500K and then 900K. - states that the NIAH harness sizes context in WORDS (~1.3x tokens) while the results table is in tokens, so the two columns are not comparable. benchmark/kimi_k3/README.md gains a pointer: all four entries there carry skip_gpu_arch gfx942, which is exactly the gap the MI300X page fills. Every tag, node count and knob value in the new page is checked against models.json and models.yaml. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…metadata The four MI300X Kimi-K3 entries could not run correctly on a cluster. Three separate defects, all from the colocated launcher inheriting conventions that only hold for the disaggregated one. Node sizing. The entries carried `"args": "-N 4 -n 4"` on the assumption that madengine forwards args to sbatch. It does not — args are appended to `bash <model>.slurm`, and neither launcher parses $@, so the flags were dropped and the job was submitted with the `slurm.nodes` default of 1. Sizing now uses `distributed.nnodes`, which madengine reconciles into `#SBATCH --nodes`, and the inert args are removed. `slurm.nodes` is deliberately left unset: setting it selects the multi-node preset, whose `NCCL_SOCKET_IFNAME=eth0` is forwarded into the container by run_xPyD_models.slurm and would override the fabric interface on an RDMA cluster. `slurm.time` is set explicitly instead, since the single-node preset's 12 h default is short for a full NIAH sweep. NIAH model tag. benchmark_niah.sh requested MODEL_PATH as the model name, correct for the disagg path where vLLM defaults served_model_name to the serve argument. serve_colocated.sh passes `--served-model-name "$MODEL_NAME"`, so every request 404'd and all three colocated entries — which default to BENCHMARK_SCRIPT=niah — would have recorded a full sweep of FAILURE rows. The harness now honors NIAH_MODEL when a launcher sets it, and the colocated launcher sets it to the tag it actually serves. Colocated run metadata. parse_to_csv derived nnodes, n_gpus, deployment_type and tags from xP/yD. serve_colocated.sh exports xP=1 yD=0 only so the shared log filenames stay unique, so a 2-node 16-GPU colocated run reported itself as a 1-node 8-GPU `disagg_1P0D` with a nixl backend it never used. Node count now comes from NNODES (exported by both launchers), and a launcher whose shape is not "xP prefill + yD decode" states its own identity via PERF_DEPLOYMENT_TYPE / PERF_TAGS. The disagg path is byte-identical to before. Also corrects the README, which documented the args-to-sbatch behavior that does not exist and a quick start that cannot work now that the "<supply-your-image>" placeholder is rejected rather than pulled. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
NUM_NODES is derived from the requested topology (xP+yD), not from the allocation, and the node list was then truncated with `head -n $NUM_NODES`. A short allocation therefore produced a silently wrong topology: pools were built from nodes that were never allocated, and the run failed much later as a connector handshake timeout with no indication of the real cause. This mattered little while these models were always launched from a matching `salloc`, but the sbatch path now sizes the allocation from the model card, so a mismatch between card and cluster is a reachable state. Check up front and name the three ways to fix it. The colocated launcher already fails on the equivalent mismatch via its TP*PP == NNODES*GPUS check, so it needs no counterpart. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ntract parse_to_csv.py hand-wrote a full 29-column perf.csv, assembling node counts, GPU counts, launcher, image and tags from the environment. Most of that is not the workload's to know: madengine already owns it, and the guesses were wrong for the colocated launcher, which sets xP=1 yD=0 only to keep log filenames unique. The NIAH writer now emits a narrow CSV — model, performance, metric, status, plus descriptive columns — and madengine merges in the run metadata via the entries' `multiple_results` declaration. The columns that genuinely describe the workload's own configuration rather than its placement (tp, pp, ep_backend, prefill_decode) move to descriptive columns, matching how scripts/vllm/run_vllm.py already reports tp/dtype/bs on the templated path. Nothing is lost; it is relocated to whichever side actually knows it. `status` stays explicit: an errored context size scores 0, and deriving status from performance would file that real failure as a SUCCESS. Scoped to leave every other workload alone. The NIAH writer is reachable only by the four Kimi entries, which all declare `multiple_results`. The throughput writer keeps its full-schema output by default, because the twelve disagg cards that use it declare no `multiple_results` and madengine reads their CSV directly with no metadata to merge — a narrow CSV there would drop every descriptive column. Their output is byte-identical; `--narrow` is the opt-in for migrating one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…uted.nnodes
madengine reads only slurm.nodes when emitting `#SBATCH --nodes`
(deployment/slurm.py: `self.nodes = slurm_config.get("nodes", 1)`), and
nothing maps distributed.nnodes across — build_orchestrator copies nnodes
into deployment_config.distributed for launcher detection only. The four
MI300X model cards carried nnodes but no slurm.nodes, so every one of them
would have been submitted as a 1-node job. The launcher's
TP*PP == NNODES*GPUS_PER_NODE assertion catches it, but only after the
allocation is granted.
Add nodes/gpus_per_node to the four slurm blocks, and ship the site configs
that carry the values a model card cannot:
- results_dir, absent from the key list madengine copies out of models.json,
and required because the slurm_multi collector globs results_dir for
perf*.csv rather than resolving multiple_results
- MODEL_DIR / LOG_PATH, which otherwise default to /shared_inference
*.json is gitignored repo-wide, so the templates need explicit negations or
they never reach a clone.
Correct the two README claims that did not match madengine's behavior
(nnodes-sizes-the-allocation, multiple_results-is-resolved-first), document
that --account/--qos are unwired for slurm_multi and need SBATCH_ACCOUNT,
and drop --keep-model-dir from the SLURM examples — it is a local-Docker
flag that madengine ignores with a warning on SLURM.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four defects, each found by running the recipes end to end on a real
MI300X SLURM cluster for the first time. None of them could have survived
a single successful run, which is consistent with these cards never having
been executed through madengine.
1. Build was not reproducible. Every source is pinned to a commit SHA, but
pip resolves build dependencies in an isolated environment from PyPI at
build time, so nothing pinned the toolchain. setuptools >= 80 added
assert isinstance(self.compiler, CCompiler)
to distutils' build_ext.build_extension, which MoRI's legacy
Cython.Distutils.build_ext path violates, so the amd_mori wheel failed
with "AssertionError: run() must precede build_extension()" while every
pinned SHA was still correct (job 223849). PIP_CONSTRAINT is the only
mechanism that reaches inside pip's build isolation -- installing
setuptools in the image does not. Set globally so the AITER, vLLM and
router stages cannot regress the same way. Build then completed in
19 minutes, all 18 stages (job 223872).
2. RDMA library mounts used -e, which is true for directories. Docker
creates a DIRECTORY at a bind-mount source that does not exist, so the
first run on a node lacking e.g. libionic.so.1 leaves an empty directory
behind, and every later run on that node mounts a directory over a file
inside the image:
OCI runtime create failed: ... not a directory
Self-propagating and per-node, so it presents as intermittent. The
surviving node then waits at socket_barrier.py forever for a peer that
already died -- 4h10m of "Waiting for nodes. . ." before the wall clock
(job 223909). Use -f so only regular files are mounted.
3. pyt_vllm_kimi-k3_mi300x_pp2xtp8 omitted the gfx942 MoE requirement that
the two wideep variants carry. gfx942 has no scaled-MXFP4 MFMA and the
a16w4 SiTUv2 heuristic FlyDSL kernel cannot codegen there, so the MoE
must be requantized to packed int4. Without it the worker aborts inside
determine_available_memory with
LLVM ERROR: Do not know how to expand this operator's operand!
and quantization_config=None in the engine config (job 224132). This is
a hardware requirement, not an EP-specific tuning, so all three
colocated cards now carry AITER_SITUV2_A8W4=1 and the int4 config.
4. A run that produced no results exited 0. The launchers warned about a
missing perf CSV and returned success anyway, so madengine recorded
exit_code=0 in its completion marker and the collector reported
"0 successful, 0 failed" -- a hard engine crash was indistinguishable
from a clean run. For a benchmark repo that is the worst possible
reporting outcome. Both launchers now exit 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Each of these cost hours to diagnose on a real cluster, and in every case the symptom is a long way from the cause -- an intermittent multi-node hang that is actually a poisoned bind-mount path, a wheel build that fails while every pinned SHA is correct, an LLVM abort that is a missing quantization flag. Record them so the next person reads instead of rediscovers. Also record the cluster facts a model card structurally cannot carry: the SLURM association wall-time cap (which silently parks a job as PENDING rather than failing), account requirements, docker on compute nodes, and the fact that a registry-less cluster has no supported path to distribute an image -- --build-on-compute requires --registry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds Kimi-K3-MXFP4 (MI300X gfx942 MXFP4 MoE) to the vllm_dissag framework: - models.yaml: K3 recipe env anchor + model entry (TP2×DP8, MoRI-EP) - run_xPyD_models.slurm: allowlists + 2P/2D validation + JIT cache split - moriio.sh: K3-gated topology (pod_hosts, api-server-count, headless KV) - vllm_disagg.sh: pod hosts export + _model_config_to_array helper - argv_assert.sh: K3 offline tests Config-only integration; image build documented in README. Co-Authored-By: Claude <noreply@anthropic.com>
…lsifiable
Both found on the first end-to-end MI300X run (job 224200), which served
Kimi-K3 and produced results -- but results that could not be trusted, from
a job that then would not exit.
Deadlock on teardown. serve_colocated.sh has the head kill only its own
server and exit, while every worker sits in `wait` on a headless vLLM that
nothing stops. srun waits on those tasks, so the job holds all nodes until
the wall clock: 4.5 hours of two idle exclusive MI300X nodes AFTER the
benchmark had already written perf.csv. The head now raises a sentinel on
the shared log volume and workers watch for it. The sentinel is raised from
an EXIT trap so a head that fails or times out cannot strand workers either.
NIAH could not distinguish truncation from retrieval failure. Scoring reads
content + reasoning_content, and K3 is a reasoning model served with
--reasoning-parser kimi_k3, so a trace that exhausts max_tokens is cut off
and only the earliest needles survive. At the old 2048 default, two of four
context sizes scored 1/10 and both listed exactly ANIMALS[0] -- the needle
placed first in the haystack. That signature is a truncated trace, not a
model failure, but finish_reason was never recorded so the two were
indistinguishable, and every row was written SUCCESS regardless.
- record finish_reason and print it as finish=<reason>
- raise the default max_tokens to 8192
- keep the needle count as the metric, always reported, never dropped
- mark a truncated row FAILURE and annotate the metric, so a measurement
artifact cannot enter the results as if it were a model result
9/10 stays SUCCESS: the README documents ~9/10 at >=20K as a known RDMA
residual, so a strict 10/10 gate would flag known-acceptable behaviour.
Logs predating the finish= field still parse.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
Adds Kimi-K3 MI300X support by integrating a new colocated multi-node vLLM launcher alongside updates to the existing disaggregated (prefill/decode) vLLM harness, plus reporting/benchmarking and documentation updates to make results consumable by madengine.
Changes:
- Introduces a 2-node colocated vLLM SLURM+container launcher (
vllm_multinode) with shared-model recipe support and shutdown coordination. - Extends the disaggregated launcher path to support
TP_SIZE>1, multi-node pool host advertisement for MoRIIO, and NIAH benchmarking/reporting. - Adds Kimi-K3 model recipes/cards, MI300X-specific docs/site-config templates, and a pinned Dockerfile for the K3 MI300X image.
Reviewed changes
Copilot reviewed 15 out of 16 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
| scripts/vllm_multinode/serve_colocated.sh | Per-node in-container entrypoint for colocated multi-node vLLM serving + benchmarking. |
| scripts/vllm_multinode/run_multinode.slurm | SLURM launcher that discovers nodes/IPs, runs per-node containers, and publishes perf CSV. |
| scripts/vllm_dissag/vllm_disagg.sh | Adds TP-aware topology math and peer-pool host export for KV connector. |
| scripts/vllm_dissag/run_xPyD_models.slurm | Adds Kimi-K3 support, TP_SIZE plumbing, NIAH option, safer RDMA mounts, perf CSV publish. |
| scripts/vllm_dissag/parse_to_csv.py | Adds narrow-schema perf CSV mode, NNODES-aware metadata, and NIAH parsing/output. |
| scripts/vllm_dissag/models.yaml | Adds Kimi-K3 MI300X serving recipe env + flags (MXFP4/int4 MoE path). |
| scripts/vllm_dissag/connectors/moriio.sh | Adds multi-node pod host list to KV config; TP-aware api-server-count and tp flag. |
| scripts/vllm_dissag/benchmark_niah.sh | Uses NIAH_MODEL when set; emits perf.csv for NIAH via parse_to_csv.py. |
| scripts/vllm_dissag/benchmark_niah.py | Raises default max_tokens, records finish_reason, and prints parseable result lines. |
| models.json | Adds Kimi-K3 MI300X disagg + colocated model-card entries with multiple_results CSV. |
| docker/pyt_vllm_kimi_k3_mi300x.ubuntu.amd.Dockerfile | New pinned build for Kimi-K3 MI300X (MoRI/AITER/vLLM/router + graft step). |
| benchmark/kimi_k3/README.md | Notes gfx942 requires multi-node sharding and points to MI300X recipes. |
| benchmark/kimi_k3/mi300x/slurm-config.disagg.json | Site-config template for 4-node disaggregated MI300X runs. |
| benchmark/kimi_k3/mi300x/slurm-config.colocated.json | Site-config template for 2-node colocated MI300X runs. |
| benchmark/kimi_k3/mi300x/README.md | Full MI300X recipe documentation, tradeoffs, and madengine integration notes. |
| .gitignore | Ensures MI300X site-config templates are not ignored. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| if [[ -n "${COLOCATED_EXTRA_ARGS:-}" ]]; then | ||
| eval "extra_args=(${COLOCATED_EXTRA_ARGS})" | ||
| serve_args+=("${extra_args[@]}") | ||
| fi |
| docker rm -f $DOCKER_CONT_NAME 2>/dev/null || true | ||
| fuser -k 2223/tcp 2>/dev/null || true | ||
| sleep 2 |
Job 239755 (Kimi-K3, 2x8 on oci-64) died in vLLM's startup barrier. Node 0 finished loading weights at 23:21:41 and entered in_the_same_node_as -> torch.distributed.barrier; node 1 was still constructing MoE layers ~50 min later. gloo gave up at 23:51:41, exactly 1800s in, and the master's TCPStore died with it, so node 1 then failed with "Broken pipe" to :29500. The barrier is all-or-nothing across all 16 ranks, so what has to stay small is the GAP between nodes, not the absolute load time. Five problems, addressed here: 1. --distributed-timeout-seconds does not reach the barrier that failed. It feeds ParallelConfig.distributed_timeout_seconds, which only configures the device/NCCL groups. The startup barrier runs on the gloo CPU group, built by GroupCoordinator from the separate cpu_distributed_timeout_seconds field. Unset, it falls back to PyTorch's stock 1800s -- which is exactly the deadline in the traceback, despite 7200 being passed. Pass --cpu-distributed-timeout-seconds too. 2. Cold AITER JIT compile on every rank, every run. The image points AITER_JIT_DIR/TRITON_CACHE_DIR/VLLM_CACHE_ROOT/COMGR_CACHE_DIR at /opt/vllm_cache, but this launcher mounted only /tmp/vllm_cache, a path nothing reads, inside --rm containers. Mount a host-persistent cache at /opt/vllm_cache keyed by image ID, as vllm_dissag/run_xPyD_models.slurm already does. Per-node cold compile is precisely the skew the barrier cannot absorb. 3. WORKER_PID was tee's PID. After `a | b &` the shell reports b, so every kill in this script hit tee and left vLLM running, and no liveness check was possible. Launch through process substitution instead. 4. A dead engine was indistinguishable from a slow one. The head polled the log for the full 4000s after its writer had already died, then reported a timeout rather than the traceback. Both wait loops now check liveness -- via a helper that also rejects zombies, since an exited-but-unreaped child still answers kill -0. This also fixes a latent hang where a worker would spin forever on a dead engine. 5. Container env was a closed allowlist. Tuning any of the above from CI meant a new -e line and another MAD PR. COLOCATED_FORWARD_ENV names the variables to forward, and they are passed via docker --env-file so values containing spaces survive. The file is node-local, mode 600, and unlike the wrapper scripts is never archived as a build artifact. Also adds an opt-out checkpoint page-cache pre-warm before the container barrier (PREWARM_CHECKPOINT=0 to skip). Its per-node duration is the real diagnostic: a large spread means the storage path is the problem and no timeout will fix it. Verified: bash -n clean on both files, and on the single-quoted srun body extracted in isolation; env-file forwarding preserves spaces and skips unset vars at mode 600; the liveness helper correctly reports an unreaped child as dead; DRY_RUN=1 emits the new flag and honours an override, with COLOCATED_EXTRA_ARGS still appended last. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
🟡 Changes recommended
The new run_multinode.slurm launcher has verified set -u-triggered failure paths (optional env expansion and hardcoded barrier port) that can prevent runs from starting or make barrier cleanup inconsistent.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
Suppressed comments (1)
scripts/vllm_multinode/run_multinode.slurm:158
- The barrier port is configurable via
CONTAINER_BARRIER_PORT(and serve_colocated.sh honors it), but thisfusercall always kills 2223. If the port is overridden, stale listeners on the configured port won’t be cleared and the barrier can hang/fail. Use the same${CONTAINER_BARRIER_PORT:-2223}default here.
fuser -k 2223/tcp 2>/dev/null || true
- Files reviewed: 15/16 changed files
- Comments generated: 1
- Review effort level: Lite
Job 240624 spent its entire 7200s allocation in the pre-warm added by the previous commit. Both nodes printed "[prewarm] reading ..." and neither ever printed "[prewarm] done", so vllm serve was never launched -- the run timed out having tested nothing. The flaw: it reads the WHOLE checkpoint on EVERY node, while PP2xTP8 means a node only ever loads its own shard. On a 1453 GiB checkpoint that is roughly 4x the necessary I/O against a single NFS export, with both nodes competing for it. It manufactured exactly the contention it was meant to relieve. Default it to off. It stays available, because the per-node duration is a useful storage probe -- a large spread between nodes means the storage path is the problem and no timeout will fix it -- but it is now opt-in via PREWARM_CHECKPOINT=1 and bounded by PREWARM_TIMEOUT_SECONDS (default 900). On expiry it warns that the filesystem could not deliver the checkpoint in the window and proceeds, rather than burning the whole allocation. Note this run did not exercise --cpu-distributed-timeout-seconds: vLLM never reached the startup barrier, so the absence of the 1800000ms gloo timeout in the log says nothing either way. That fix is still untested. Verified: default unset is silent; PREWARM_CHECKPOINT=1 runs and reports its duration; a stalled read hits the budget, warns, and exits 0 so startup continues. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both GLM-5.1 cards passed at 1P/1D with each pool running TP8 and one DP rank: cluster.sh exports TP_SIZE=GPUS_PER_NODE, and the wideEP path read TP_SIZE. When the wideEP layout moved to its own knob, EP_TP_SIZE (default 1), these cards silently became TP1 x 8 DP ranks. Every GPU then holds a full copy of the non-expert weights, and decode crashes in MoRI EP setup with the GPUs nearly full. EP_TP_SIZE=8 on the cards restores the layout they passed with. The serve argv matches the earlier one: -tp 8 is spelled --tensor-parallel-size 8, and the pod-host list and --moriio-dp-size that EP_TP_SIZE>1 adds are one host and 1 at 1P/1D, which is what the earlier path addressed. The recipe is unchanged.
The check reads each card's own env_vars, so a card that loses EP_TP_SIZE fails here: TP8, one DP rank per pool, and the single peer host in the pod-host list, for both roles. It fails on the cards as they were before.
pyt_vllm_disagg_deepep_deepseek-v3 still failed after 52ab5e1 gave it the recipe's block size: decode logged "Memory access fault by GPU" on every GPU while capturing its first decode cudagraph, right after loading AITER's MLA kernel (mla_a8w8_qh128...), and the workers died. The DeepSeek recipe turns AITER MLA off (VLLM_ROCM_USE_AITER_MLA=0), for that reason, and asks for PIECEWISE decode cudagraphs. The deepep setup_env exported VLLM_ROCM_USE_AITER_MLA=1 over it, and the decode launch hardcoded FULL_DECODE_ONLY, so decode ran the ROCM_AITER_MLA backend in full-graph mode. The moriio connector honours the recipe, picks TRITON_MLA with PIECEWISE on the same model, and passes. The same lapse as 52ab5e1 and 4a28f2c. setup_env now takes the AITER knobs the recipes set from the environment, and decode takes DECODE_CUDAGRAPH_MODE (then VLLM_CUDAGRAPH_MODE) and CUDAGRAPH_CAPTURE_SIZES, each defaulting to the value that was hardcoded, so a recipe that sets none of them is unchanged. This also covers pyt_vllm_disagg_nixl_deepseek-v3, which runs the same deepep path. The dry run now prints the server's AITER env after the argv block, so argv_assert can check it. argv_assert: 143 passed, 0 failed; four of the new checks fail against the previous rixl.sh. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pyt_vllm_disagg_mori_kimi-k3 failed prefill start-up with CUDA out of memory on every GPU while allocating the KV cache. Each GPU had 172.6 GiB free before loading (the 16 GiB MoRI heap already taken), the TP2 weights took 148.56 GiB, and KV_CACHE_MEMORY_BYTES=40e9 then asked for 37.25 GiB of the ~24 GiB left. The recipe comment's budget assumed 137.5 GiB of weights, which even then left no headroom. 16e9 (~1.13M tokens) leaves ~9 GiB for activations, decode cudagraphs and RCCL, and still covers --max-model-len 1000000. Kimi-K3-MXFP4 passes NIAH up to 280K tokens at the same TP2 x DP8 layout with 8e9. Single requests beyond roughly 250K tokens may no longer fit; the card's NIAH sweep stays within that. The MI355X recipe keeps 40e9. argv_assert: 145 passed, 0 failed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…that failed parse_to_csv.py looked for a result's [RUNNING] header in the text since the previous result and took the first one there. When a cell stalls it prints no result block, so the next cell's text holds both headers: on pyt_vllm_disagg_deepep_deepseek-v3 the con=256 result (half its requests failed) was filed as con=128, and the con=256 row vanished. It also never looked before the first result, so the first measured cell of every sweep was dropped: a passing DeepSeek-V3 MoRI run lost its con=8 row (512.27 tok/s). Every sweep row was also written as SUCCESS. That run's con=512 cell failed all 1024 requests and was reported as "SUCCESS 0.00 tok/s", and the build passed. The parser now takes the last header before each result, including the first, records failed requests, and gives a [STALL] cell a zero-throughput row. A cell that stalled, lost any request, or measured no throughput is a FAILURE row, as agentic rows already are. The CONCURRENCY.csv keeps its columns. tests/parse_to_csv_assert.sh: 7 passed, 0 failed; 5 of them fail against the previous parser. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rker OOM With the KV cache fitting, pyt_vllm_disagg_mori_kimi-k3 got its prefill master up and then lost decode: one decode worker ran out of GPU memory on its first cudagraph capture (PIECEWISE, the largest size), when AITER's fused-MoE asked for a 3 GiB workspace with ~9 GiB left after the 148.56 GiB of weights and 14.9 GiB of KV cache. Decode captured sizes up to 256 while the recipe runs --max-num-seqs 8, so every size above 8 was memory it could never use. The recipe now sets CUDAGRAPH_CAPTURE_SIZES="1 2 4 8". The OOM did not stop the job: the other workers waited, the engine never printed an initialization failure, and the job sat ~55 min until vLLM's own 3600s engine-ready timeout. The start-up watch now treats a worker's torch.OutOfMemoryError as fatal, as it already does an NCCL or HIP failure. argv_assert: 149 passed, 0 failed; two of the new checks fail against the previous recipe and launcher. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Reverts the recipe side of 8370e44 (KV_CACHE_MEMORY_BYTES 40e9 -> 16e9) and b3c5702 (CUDAGRAPH_CAPTURE_SIZES "1 2 4 8"). Both made the MI300X Kimi-K3 disagg card fit in memory by giving up what the recipe asks for: the 2.84M-token KV cache for single requests beyond ~600K tokens, and its decode cudagraph sizes. The recipe is restored as it was; the out-of-memory it hits on MI300X is to be fixed, not traded away. What the two failed runs measured, for whoever takes that on: 148.56 GiB of weights per GPU at TP2 (the recipe comment budgets 137.5 GiB), 172.6 GiB free before loading with the 16 GiB MoRI heap. The same image also logs "This AITER build ignores the SiTUv2 activation on the packed-int4 MoE path and would silently compute SiLU" (needs ROCm/aiter#4471). Kept from b3c5702: the start-up watch treats a worker's torch.OutOfMemoryError as fatal. argv_assert: 145 passed, 0 failed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pyt_vllm_disagg_mori_kimi-k3-mxfp4 came up on all four nodes, prefill and decode, and then failed every NIAH request with HTTP 500: the router's prefill call got 404 "The model `<path>` does not exist." The recipe (the PR 241-validated one) starts vLLM with --served-model-name kimi-k3, while benchmark_niah.sh requested MODEL_PATH, on the assumption that disagg recipes never set a served name. vllm_disagg.sh now resolves SERVED_MODEL_NAME from the recipe's own flags (--served-model-name when present, else MODEL_PATH, vLLM's default), and NIAH requests it unless NIAH_MODEL is set. It reads the recipe rather than the router's /v1/models, which 503s under MoRIIO service discovery. No recipe changes; every recipe without --served-model-name resolves to MODEL_PATH exactly as before. The dry run now also prints SERVED_MODEL_NAME. argv_assert: 149 passed, 0 failed; the four new checks fail against the previous launcher. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pyt_vllm_kimi_k3_mi300x, pyt_vllm_kimi_k3_mi355x and vllm_disagg_inference.kimik3 were one stack three times. vLLM (862bfd8, the head of the branch the disagg file named), MoRI (v1.2.2) and AITER were identical in all three; they differed in the GPU arch, the router pin (the disagg file's was 27 commits newer, and colocated cards do not use the router), and NIXL on or off. Each also built on the ROCm vLLM CI base, installed AITER 0.1.19, then deleted it and copied aiter, aiter_meta and flydsl out of a second, "proven" image. Both proven images (amdsiloai/vllm:kimi-k3-mi325x-release-v2 and vllm/vllm-openai-rocm:kimi-k3) carry the same layer, digest e2b3951e36ca: a checkout of carlushuang/aiter-k3 branch k3-for-amd at ROCm/aiter 68e42f5f (0.1.17.dev395), with the Kimi-K3 tuned MoE configs and aiter.ops.triton.conv, and flydsl 0.2.4. (The old header's "the donor's flydsl is 0.2.2" was wrong.) docker/vllm_kimi_k3.ubuntu.amd.Dockerfile builds that AITER commit from source, the same way the GLM-5.1 image builds its AITER, so there is one base and no second image. It builds for MAD_SYSTEM_GPU_ARCHITECTURE (gfx942 or gfx950, no default: a wrong-arch image fails at runtime, so the build refuses to guess), the build arg madengine already passes; the arch reaches each step as K3_GFX_ARCH so madengine's Dockerfile arch check does not read a fixed arch. The router takes the disagg pin by SHA (82dc981), rocSHMEM its pin, and WITH_NIXL defaults on so the image serves every connector. Checks read package metadata only: importing aiter probes the GPU, and the build agent has none. A prefix no other Dockerfile shares, so madengine's `<prefix>.*` lookup finds only this file even without its exact-match fix. All seven Kimi-K3 cards point at it; every card in scripts/*/models.json resolves to exactly one Dockerfile. The MI355X bring-up caveats move into its STATUS. argv_assert: 149 passed; parse_to_csv_assert: 7 passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Kimi-K3 image builds for MAD_SYSTEM_GPU_ARCHITECTURE and refuses to guess; madengine cannot detect it on a build host without a GPU. The m2m profile now states gfx942 as docker_build_arg, so a build with no probed or given arch still gets the cluster's. An arch the CI probes or is given overrides it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sets scripts/common/clusters/m2m.json held partition amd-rccl, 8 GPUs per node and exclusive, which are madengine's own SLURM presets and are also set in cluster.sh. Nothing in MAD read it; only the MAD CI resolver did. The build arch added to it in 08c45d3 moves to the CI's own target table (rocAutomation NODE_PROFILES: oci-64 -> gfx942), which is where knowledge of CI targets belongs. Removed: clusters/m2m.json, clusters/README.md and the .gitignore exception for them. cluster.sh and scripts/common/README.md now point at madengine's presets and --additional-context instead of the profile. The allocation a run gets is unchanged: the CI's path-equivalence check resolves all 46 multinode cards to --partition=amd-rccl --gpus-per-node=8 --exclusive, the values the profile supplied. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
run_xPyD_models.slurm ended with the container-cleanup srun, so its exit status, and the SLURM job's, was the cleanup's. A run whose prefill died at start-up (pyt_sglang_disagg_mori_io_agentic_qwen3-32b, GPUs already occupied on a node) ended COMPLETED with no results, and madengine collected "0 successful, 0 failed" and reported success. The job now fails when the main srun exits non-zero (a node's container exits non-zero when its server or benchmark fails; the passing runs checked exited 0 on every node) or when no /run_logs/$SLURM_JOB_ID/perf.csv was written, as the vLLM launchers already require. SKIP_BENCHMARK=1 serves without benchmarking and is not required to produce results. Checked with a stubbed srun: clean run 0; node failure 1; no perf.csv 1; SKIP_BENCHMARK=1 without perf.csv 0; SKIP_BENCHMARK=1 with a node failure 2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
parse_to_csv.py took each result block's concurrency from the header before it and wrote every row as SUCCESS. On pyt_sglang_disagg_mori_io_qwen3-32b the con=256 cell lost half its requests to connection errors (256 of 512 succeeded) and was recorded SUCCESS at 6390 tok/s, and the con=512 cell, whose warmup got "Bad Gateway" and printed no result, had no row at all; the build passed. Each RUNNING line now owns the text up to the next one. A cell that printed no result is a FAILURE row at its own concurrency; sglang.bench_serving prints no failed count, so a cell whose successful requests fall short of its prompts is FAILURE; zero throughput is FAILURE. Clean runs are unchanged: re-parsing the archived logs of two passing runs gives the same 7 SUCCESS rows and numbers. tests/parse_to_csv_assert.sh (new): 7 passed; 4 fail against the previous parser. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d, and publish long_context results
Three gaps in how vLLM disagg results become perf.csv rows, each of which let a
broken run pass or a working one fail:
- A sweep cell whose benchmark crashed without the harness's [STALL] line (for
example a warmup that got an error from a dead server) printed no result and
got no row. parse_to_csv.py now reads the log cell by cell, as the SGLang
parser does: each [RUNNING] line owns the text up to the next one, and a cell
with no result is a FAILURE row.
- NIAH: a context length whose every request errored prints NO-RESULT and got no
row, and a length that lost some requests ("[k timeout/err excluded]") was
SUCCESS on the survivors' mean. Both are now FAILURE rows; a low or truncated
retrieval score is still a measurement and stays SUCCESS. madengine counts
FAILURE rows as failed runs and exits non-zero, so an all-errored NIAH run
still fails, now with a row per length saying which.
- benchmark_long_context.sh never wrote perf.csv, so a long_context run always
ended "no perf CSV" and failed. It now calls the parser as the sweep does,
and the parser reads its "[RUNNING] isl=... con=..." header form.
Re-parsing archived logs: three passing sweeps and three passing NIAH runs are
unchanged; the all-NO-RESULT Kimi-K3-MXFP4 run gives four FAILURE rows. The
long_context script, run end to end with a stubbed vllm, writes a SUCCESS row
for a completed cell and a FAILURE row for a timed-out one.
tests/parse_to_csv_assert.sh: 16 passed; the new cases fail against the
previous parser.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…extending MAD
MAD had no docs folder; what a reader needed was spread over the root README and
READMEs nested under scripts/ and benchmark/. docs/ now holds it in one place, in
depth, in the order someone new needs it:
- README.md: what MAD is, a learning path, a glossary, and an index (with the
madengine docs it builds on)
- getting-started.md, adding-a-model.md: running models; every card field,
Dockerfile matching, MAD_SYSTEM_GPU_ARCHITECTURE
- multinode-overview.md, multinode-running.md: disaggregated and colocated
inference, launchers, connectors, topology, the architecture diagrams; running
through madengine, sbatch and salloc, weights, logs, failure modes
- configuration.md: every configuration layer, what wins, and where to change what
- benchmarks-and-results.md: each benchmark, perf.csv rows and what makes a run pass
- vllm-disagg.md, sglang-disagg.md, kimi-k3.md: full references
Where a README disagrees with the code, the pages follow the code. The existing
READMEs stay. The two this PR added, scripts/common/README.md and
benchmark/kimi_k3/mi300x/README.md, now point into docs/, and the root README links
to docs/README.md.
benchmark/kimi_k3/mi300x/slurm-config.{colocated,disagg}.json are removed, with
their .gitignore exception. They were fill-in templates that restated values with a
source elsewhere: the partition, GPUs per node, exclusivity and output directory
are madengine's SLURM presets, the node count is each card's, and MODEL_DIR and
LOG_PATH default in scripts/common/cluster.sh. The docs pass only what differs, as
--additional-context, and say where every other value comes from.
The Kimi-K3 Dockerfile's build note no longer says the context must be the repo
root: the image copies nothing from the build context.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… have it Both Kimi-K3 image builds failed at the AITER step with "fatal: reference is not a tree: 68e42f5f". The release images built AITER from branch k3-for-amd of a fork that is not public; the commit is in ROCm/aiter's fork network (GitHub serves it by SHA and its API finds it) but on none of ROCm/aiter's branches, so git clone never downloads it. The step now fetches that exact commit by SHA, with --tags for the release tags its version is derived from, and checks it out. Checked outside docker: HEAD 68e42f5f4, the composable_kernel submodule, the six kimik3 tuned configs and aiter/ops/triton/conv are all present. The version string is now derived from v0.1.18 (tags added upstream since the release images were built) instead of 0.1.17.dev395; the code is the same commit, and the image's check requires only that the version names 68e42f5f. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(built on top of existing kimi k3 pr for madengine enablement and testing)
Summary
Adds Kimi-K3 multinode inference to MAD and makes the multinode vLLM / SGLang workloads run, fail and report the same way whether they are launched through madengine or submitted directly with
sbatch. It also adds adocs/folder that takes a new reader from zero to running, configuring and extending these workloads.Start here:
docs/README.md→docs/multinode-running.md.Depends on: ROCm/madengine#203. It is needed for exact Dockerfile matching, for counting a failed SLURM deployment as a failure, and for the
slurm_multicollector reading a card'smultiple_results.What's in this PR
Kimi-K3 on MI300X / MI355X: 7 new cards, 1 image
pyt_vllm_kimi-k3_mi300x_wideep_moriepscripts/vllm_multinode(colocated)pyt_vllm_kimi-k3_mi300x_wideep_allgatherpyt_vllm_kimi-k3_mi300x_pp2xtp8pyt_vllm_kimi-k3_mi300x_pp2xtp8_way4mad-config.kimi-k3.yamlpyt_vllm_disagg_mori_kimi-k3-mxfp4scripts/vllm_dissag(disaggregated)pyt_vllm_disagg_mori_kimi-k3pyt_vllm_disagg_mori_kimi-k3-mxfp4_mi355xdocker/vllm_kimi_k3.ubuntu.amd.Dockerfile.MAD_SYSTEM_GPU_ARCHITECTURE(gfx942orgfx950, no default).68e42f5f), with flydsl 0.2.4. Nothing is copied from a second image.REQUIRE_LOCAL_WEIGHTS=1on the Kimi cards requires local NVMe copies.MODEL_WEIGHTS_NAMEselects the checkpoint directory when it is staged under another name (for exampleKimi-K3vsKimi-K3-MXFP4).Launchers: every failure ends the job with the reason in the log
perf.csv.--served-model-name.Results:
perf.csvsays what happened, cell by cellFAILURErow. It is no longer a missing row or a falseSUCCESS.NO-RESULT, is aFAILURErow.long_contextruns now writeperf.csv.Configuration: one source for each setting
scripts/common/cluster.sh.models.yamlrecipe.--additional-context.docs/configuration.mdlists every layer, what wins, and where to change what.How to run
Install madengine with the changes this PR depends on:
Through madengine
From the repo root on a SLURM login node. The Kimi cards carry
DOCKER_IMAGE_NAME: "<supply-your-image>", so either build and push an image or point at one.What you usually need to pass: only the time limit, because the cards say 24:00:00 and most partitions cap lower.
Where the rest comes from: the node count is in each card. The partition, GPUs per node and exclusivity are madengine's presets. Model and log paths default in
scripts/common/cluster.sh.On a cluster that differs, add only the keys that differ:
{"slurm": {"partition": "<partition>", "time": "06:00:00"}, "env_vars": {"MODEL_DIR": "/path/to/models", "LOG_PATH": "/path/to/shared/logs"}}If a node stages the checkpoint under another name, add
"MODEL_WEIGHTS_NAME": "Kimi-K3"underenv_vars.Standalone with
sbatchEach launcher takes the card's
env_varsas its contract. Export them, name the image, and submit with the allocation.sbatchoptions go before the script.$LOG_PATH/<job>/perf.csv, one row per cell withSUCCESSorFAILURE;docs/multinode-running.md(inside ansalloc, weights, fabric, logs, fail-fast) anddocs/kimi-k3.md.Checks that need no GPUs
Validation status (MI300X, gfx942)
Known issues
pyt_vllm_disagg_deepep_deepseek-v3,pyt_vllm_disagg_nixl_deepseek-v3):DeepEP error: CPU recv timeoutin high-throughput intranode dispatch;FAILURE.pyt_vllm_disagg_mori_kimi-k3(theKimi-K3recipe): its memory budget (40e9 KV cache and a 16 GiB MoRI heap on top of ~148.6 GiB of TP2 weights) does not fit a 192 GB MI300X. It is left unchanged pending the recipe owner.pyt_vllm_disagg_mori_kimi-k3-mxfp4is the validated Kimi disagg recipe.Review guide
docker/vllm_kimi_k3.ubuntu.amd.Dockerfile: the single Kimi image.scripts/vllm_dissag/,scripts/sglang_disagg/,scripts/vllm_multinode/: launchers, recipes (models.yaml), cards (models.json), parsers.scripts/common/cluster.sh: site facts shared by all three launchers.docs/: new. The existing READMEs are unchanged except where they now point intodocs/.